Papers with sequences exceeding
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling (2025.findings-emnlp)
Copied to clipboard
Zeyu Liu, Souvik Kundu, Lianghao Jiang, Anni Li, Srikanth Ronanki, Sravan Babu Bodapati, Gourav Datta, Peter Anthony Beerel
| Challenge: | a novel linearization framework is proposed to reduce the cost of training transformers from scratch. |
| Approach: | They propose a linear attention framework that integrates pre-trained transformers into a performant linear attention architecture. |
| Outcome: | The proposed framework improves performance on mistral-7B with 1K-length sequences and BABILong benchmarks. |